Papers with human evaluation scores

10 papers
Neural Machine Translation System using a Content-equivalently Translated Parallel Corpus for the Newswire Translation Tasks at WAT 2019 (D19-52)

Copied to clipboard

Challenge: In addition to the JIJI Corpus, we developed a corpus of 0.22M sentence pairs by manually, translating Japanese news sentences into English content- equivalently.
Approach: They propose to use JIJI Corpus and Equivalent-style sentences to translate Japanese news sentences into English content- equivalently.
Outcome: The proposed translation models achieved the best human evaluation scores in the newswire translation tasks at WAT 2019 . they used the JIJI Corpus, which was provided by the task organizer, and the Equivalent-style translation model to translate Japanese news sentences into English content- equivalently.
Entity-Based Semantic Adequacy for Data-to-Text Generation (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing pre-trained language models have improved the fluency of text generation systems, but semantic adequacy remains an unsolved issue.
Approach: They propose an automatic evaluation metric to assess to what extent models that verbalise RDF graphs produce text that contains mentions of entities occurring in the input.
Outcome: The proposed metric can be used to assess to what extent generation models verbalise RDF graphs produce text that contains mentions of the entities occurring in the input.
RoViST: Learning Robust Metrics for Visual Storytelling (2022.findings-naacl)

Copied to clipboard

Challenge: Visual storytelling is the task of generating a story paragraph that describes a given image sequence.
Approach: They propose 3 evaluation metrics sets that analyze which aspects we would look for in a good story . they compare their correlation with human judgement scores on a sample of machine stories .
Outcome: The proposed evaluation metrics outperform other metrics on human correlation on a sample of machine stories from state-of-the-art models.
A Deep Analysis of the Impact of Multiword Expressions and Named Entities on Chinese-English Machine Translations (2024.findings-emnlp)

Copied to clipboard

Challenge: a study on the impact of multiword expressions and multiword named entities (NEs) on the performance of Chinese-English machine translation systems is presented.
Approach: They propose to use Chinese multiword expressions and multiword named entities (NEs) to evaluate machine translation performance.
Outcome: The proposed methods show that Chinese-English machine translation systems perform significantly worse on Chinese sentences with most kinds of MWEs and NEs.
Generating Classical Chinese Poems via Conditional Variational Autoencoder and Adversarial Training (D18-1)

Copied to clipboard

Challenge: Existing models for automatic poetry generation lack term novelty and thematic consistency.
Approach: They propose a conditional variational autoencoder with adversarial training for classical Chinese poem generation.
Outcome: The proposed model outperforms existing models on a large poetry corpus on 'classical Chinese' . it generates poems with novel terms and learns their thematic consistency with their titles.
A Large-Scale Study of Machine Translation in Turkic Languages (2021.emnlp-main)

Copied to clipboard

Challenge: a large corpus covering 22 Turkic languages is included in this paper . low-resource MT evaluation has traditionally focused on European languages due to limitations of available technology and resources.
Approach: They present a case study of the practical application of MT in the Turkic language family . they propose to realize the gains of NMT for Turkic languages under high-resource to extremely low-resourced scenarios.
Outcome: The proposed study shows that the new methods can be used in the Turkic language family . the results highlight bottlenecks in building competitive systems .
LCFO: Long Context and Long Form Output Dataset and Benchmarking (2025.findings-acl)

Copied to clipboard

Challenge: Using long text outputs to evaluate progress in summarization and summary expansion tasks is challenging.
Approach: They propose a framework for assessing gradual summarization and summary expansion capabilities across diverse domains.
Outcome: The proposed framework provides alignments between specific QA pairs and corresponding summaries in 7 domains.
Perturbation CheckLists for Evaluating NLG Evaluation Metrics (2021.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics for natural language generation are inadequate . existing metrics are not robust against simple perturbations and disagree with scores assigned by humans to perturbed output.
Approach: They propose to propose checks which perturb the output and target a specific criteria and then use them to refine their evaluation.
Outcome: The proposed templates show that existing evaluation metrics are not robust against simple perturbations and disagree with human scores on the perturbed output.
Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation Approach (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for evaluation of dialog systems are expensive and not scalable . a framework for estimating human evaluation scores is proposed to bridge this gap .
Approach: They propose a framework for estimating human evaluation scores based on off-policy evaluation . they use language quality metrics for single-turn response generation given a fixed context .
Outcome: The proposed framework outperforms existing methods in terms of correlation with human evaluation scores.
Translationese as a Language in “Multilingual” NMT (2020.acl-main)

Copied to clipboard

Challenge: Recent work examines the impact of translationese in machine translation evaluation using the WMT evaluation campaign.
Approach: They propose to use a sentence-level classifier to distinguish translationese from original target text to generate a machine translation model that can produce more natural outputs at test time.
Outcome: The proposed model produces more natural outputs at test time, yielding gains in human evaluation scores on accuracy and fluency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations